Papers with annotation procedures

5 papers
Korean-Specific Emotion Annotation Procedure Using N-Gram-Based Distant Supervision and Korean-Specific-Feature-Based Distant Supervision (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to annotate unlabeled data with emotions are expensive and time-consuming.
Approach: They propose an annotation procedure that leverages Korean emotion lexicons and Korean-specific emotion features to annotate unlabeled data.
Outcome: The proposed procedure compares with the KTEA dataset and a large-scale emotion-labeled dataset.
Efficient Pairwise Annotation of Argument Quality (2020.acl-main)

Copied to clipboard

Challenge: Especially crowdsourcing suffers from assessors having different reference frames to base their judgments on and task instructions being nondescript and therefore unhelpful in ensuring consistency.
Approach: They propose an efficient annotation framework for argument quality that uses a stochastic transitivity model and an effective sampling strategy to infer high-quality labels.
Outcome: The proposed model significantly outperforms existing annotation procedures and offers statistical insights into argument quality.
Polish Discourse Corpus (PDC): Corpus Design, ISO-Compliant Annotation, Data Highlights, and Parser Development (2024.lrec-main)

Copied to clipboard

Challenge: The Polish Discourse Corpus employs ISO 24617-8 for discourse relation annotation.
Approach: They propose to adopt ISO 24617-8 standard for discourse relation annotation for Polish and to develop a parser tailored for the framework.
Outcome: The Polish Discourse Corpus adopts ISO 24617-8, a segment of the Language Resource Management – Semantic Annotation Framework (SemAF) the paper examines the corpus architecture, annotation procedures, and the challenges encountered by annotators.
RAAMove: A Corpus for Analyzing Moves in Research Article Abstracts (2024.lrec-main)

Copied to clipboard

Challenge: RAAMove is a comprehensive multi-domain corpus dedicated to the annotation of move structures in Research Article (RA) abstracts.
Approach: They propose a multi-domain corpus dedicated to the annotation of move structures in RA abstracts.
Outcome: The proposed corpus is based on a human-annotated dataset and a BERT-based model to verify its effectiveness.
SM-FEEL-BG - the First Bulgarian Datasets and Classifiers for Detecting Feelings, Emotions, and Sentiments of Bulgarian Social Media Text (2024.lrec-main)

Copied to clipboard

Challenge: SM-FEEL-BG is the first Bulgarian-language package for emotion detection and sentiment analysis.
Approach: They introduce SM-FEEL-BG, a Bulgarian-language package that contains 6 datasets with Social Media (SM) texts with emotion, feeling, and sentiment labels and 4 classifiers trained on them.
Outcome: The proposed package is the first to be released in Bulgarian and is available for free.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations